Back

Medical Physics

Wiley

Preprints posted in the last 90 days, ranked by how well they match Medical Physics's content profile, based on 14 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
TCIA Radiology Image Processing for AI and Radiomics

Rich, J. M.; Kang, R.; Jin, D.; Subramanian, S.; Duddalwar, V.; Pachter, L.

2026-06-24 radiology and imaging 10.64898/2026.06.15.26354651 medRxiv
Top 0.1%
28.4%
Show abstract

We developed a standardized, reproducible preprocessing framework for computed tomography (CT) imaging data from multi-institutional repositories such The Cancer Imaging Archive (TCIA), enabling consistent radiomics and artificial intelligence (AI) analyses. Imaging data from TCGA-KIRC patients available on TCIA were used as a representative heterogeneous dataset characterized by variation in acquisition protocols, inconsistent metadata, and differing image quality. The proposed modular pipeline includes series filtering, DICOM-to-NIfTI conversion, orientation harmonization to a canonical coordinate system, voxel spacing normalization, intensity clipping and normalization, segmentation integration, and metadata validation, and is implemented in a reproducible, notebook-based framework compatible with common radiomics and deep learning workflows. This pipeline standardizes imaging data into analysis-ready volumes with consistent geometry, intensity distributions, and spatial alignment, reducing non-biological variability that can adversely affect radiomic feature stability and model performance. The modular design enables task-specific adaptation of individual preprocessing steps while maintaining overall consistency. Although demonstrated on TCIA, this framework is generalizable to other heterogeneous imaging datasets and provides a foundation for robust, large-scale computational imaging studies.

2
A Real-World Evaluation of Failure Detection for Liver CT Segmentation

Bennett, J.; Woodland, M.; Castelo, A.; Altaie, M.; Antony, A.; Siddiqi, N. S.; Long, J. P.; Brock, K. K.

2026-06-29 radiology and imaging 10.64898/2026.06.26.26356692 medRxiv
Top 0.1%
19.3%
Show abstract

Deep learning models deployed in clinical imaging frequently encounter distribution shifts, yet most out-of-distribution (OOD) detection methods are evaluated only on controlled research datasets. As a result, it is unclear whether existing approaches can reliably identify segmentation failures that arise in real-world clinical practice. We evaluated six OOD detection methods on a deployed liver CT segmentation model (3D nnU-Net) using internal data from 400 patients and external data from 100 patients collected across nearly 70 sites in 7 countries. One method was Pairwise Surface DSC, a surface-based extension of Pairwise DSC, that we introduced. OOD performance was measured using sensitivity, AUROC, and balanced accuracy, with thresholds determined on an independent cohort of 400 patients using the Youden J statistic. Statistical significance was assessed using McNemar tests and stratified bootstraps ( = 0.05) with Benjamini-Hochberg correction. Pairwise Surface DSC was the top-performing method, with perfect sensitivities (1.00), near-perfect AUROCs (0.97 internal; 1.00 external), and the highest balanced accuracies (0.94 internal; 0.88 external; p<0.001). These results show that automated failure detection for liver CT segmentation is clinically feasible and that Pairwise Surface DSC is a promising candidate for deployment. Our code is available at https://github.com/mckellwoodland/liver_ct_ood_translation.

3
Dual-Filament 3D Printing of Patient-Specific CT Phantoms with Embedded Implants and Tunable Metal-Artifact Intensity

Pasyar, P.; Mei, K.; Im, J. Y.; Roshkovan, L.; Geagan, M.; Noël, P. B.

2026-07-20 radiology and imaging 10.64898/2026.07.17.26358319 medRxiv
Top 0.1%
18.9%
Show abstract

ABSTRACT Background: Metallic implants such as orthopedic screws, prostheses, and dental hardware produce beam-hardening, photon-starvation, and streak artifacts that degrade computed tomography (CT) image quality, and the metal artifact reduction (MAR) methods developed to mitigate them require objective, reproducible benchmarking. Purpose: Objective evaluation of MAR algorithms in CT is hindered by the absence of phantoms that simultaneously provide anatomically realistic backgrounds, embedded implants of known geometry, and controllable, ground-truth--referenced artifact intensity. We present a dual-filament, voxel-level three-dimensional (3D) printing method that fulfills these requirements and demonstrate its capabilities on a clinically representative cervical spine case with embedded orthopedic spinal screws. Methods: The proposed method extends the PixelPrint framework, a fused-deposition-modeling (FDM) workflow that converts clinical Digital Imaging and Communications in Medicine (DICOM) data directly into 3D-printer Geometric code (G-code) without intermediate segmentation or surface meshing, to interleaved, voxel-level deposition of two filaments: a calcium-doped polylactic acid (PLA) for soft tissue and bone, and a higher-attenuation metal-doped PLA for metallic implants. For demonstration, anonymized DICOM data of a healthy cervical spine were used to design and fabricate three matched phantoms, each with six embedded spinal screws at C4--C6: a 0% metal-infill ground-truth phantom, a 50% medium-metal-infill phantom, and an 85% high-metal-infill phantom. All phantoms were scanned on a clinical spectral CT system at 120 kVp and 1000 mAs, reconstructed at 0.67 mm slice thickness with virtual monoenergetic imaging (VMI) across 50--190 keV. Method performance was characterized by region of interest (ROI)-based Hounsfield Unit (HU) agreement with the source patient data and by the noise-independent Gumbel-distribution p-index metric. Results: The dual-filament method reproduced patient anatomy, soft-tissue contrast, and screw geometry with high fidelity. ROI HU values agreed with patient data within {+/-}25 HU for soft tissue and trabecular bone; cortical regions were underestimated owing to the current ceiling of the calcium-doped PLA used in this study. The tunable-artifact behavior was quantified as follows: the Gumbel location parameter scaled monotonically from 46.7 HU (no-metal background) to 57.1 HU (50% infill) to 90.5 HU (85% infill) for the VMI 70 keV with standard filter. High-keV VMI reconstructions substantially reduced streak and beam-hardening artifacts while preserving anatomic detail. Conclusions: The proposed dual-filament, voxel-level PixelPrint method enables the fabrication of patient-specific, multi-material CT phantoms with embedded metallic implants and controllable, ground-truth--referenced artifact intensity. Although demonstrated here in a single cervical-spine case, the workflow is anatomy- and implant-agnostic by construction and could in principle be adapted to other musculoskeletal sites (e.g., knee, hip, dental) and implant materials, providing a reproducible methodological foundation for benchmarking MAR algorithms, characterizing spectral CT performance, and validating emerging photon-counting detector systems. Keywords: 3D printing methodology; fused deposition modeling; voxel-level multi-material printing; spectral computed tomography; metal artifact reduction; phantom design; orthopedic implants; dual filament; PixelPrint.

4
BioDeformUNet: A Deep Learning Model for Biomechanically Informed Liver Image Registration

Zhang, X.; OConnor, C.; Castelo, A.; Woodland, M.; Daoud, B.; Paolucci, I.; Albuquerque, J.; Altaie, M. A.; Siddiqi, N.; Patel, A.; Odisio, B.; Brock, K.

2026-08-05 radiology and imaging 10.64898/2026.08.03.26359612 medRxiv
Top 0.1%
18.6%
Show abstract

Purpose: To build a 3D U-Net model, BioDeformUNet, to predict the deformation vector field (DVF) of the liver in near real-time, for efficient intra-procedural evaluation of the minimal ablative margin (MAM). Materials and Methods: This retrospective study included 170 contrast-enhanced computed tomography (CECT) image pairs from 157 patients who underwent liver ablation treatment between 2020-2024. Each data instance included one pre-ablation CECT (pre-CECT) and one post-ablation CECT (post-CECT). BioDeformUNet was trained under the guidance of DVFs generated by a biomechanical model-based deformable image registration (DIR) algorithm using a loss function that focused on large liver deformations. Data were split patient-wise into training (92-93 patients), validation (23-24 patients), and testing sets (42 patients). We compared our performance with two deep learning-based DIR methods: VoxelMorph and VFA. Evaluation metrics included: target registration error (TRE), Dice similarity coefficient (DSC), Minimum Ablation Margin (MAM), and inference time. For BioDeformUNet, we additionally evaluated the accuracy of the deformed tumor center-of-mass mapping by comparing the predicted tumor center location with that generated by Morfeus. A mapping error less than 3.0 mm (corresponding to the voxel size) was considered accurate. We used the Wilcoxon signed-rank test to assess the significancy of each test result. Our code is available at https://github.com/XinyueZhang831/BioDeformUNET. Results: The TRE of BioDeformUNet was not significantly different from Morfeus (3.31 BioDeformUNet; 3.23 Morfeus; p-value=0.41). The BioDeformUNet DVF magnitude was within 3.0 mm of Morfeus DVF for an average of 91.9% of the voxels. Tumor mapping errors greater than 3.0 mm occurred in only 8 cases. The inference time of BioDeformUNet was 0.6s per image pair, 0.2s for VoxelMorph, 0.3s for VFA, and 20.2s for Morfeus. Conclusion: BioDeformUNet achieved a similar performance to the biomechanical model-based algorithm but required fewer computational operations, resulting in a 34 times speedup in DVF computation.

5
Artificial Intelligence-Enabled Detection of Vascular Perfusion Defects on Ventilation/Perfusion (V/Q) Scintigraphy for Pulmonary Embolism

Jabbarpour, A.; Moulton, E.; Kaviani, S.; Zeng, W.; Ghassel, S.; Akbarian, R.; Couture, A.; Roy, A.; Liu, R.; Al-ali, Y.; Foufa, Y.; Hejji, N.; AlSulaiman, S.; Shirazi, Z.; Leung, E.; Klein, R.

2026-07-08 radiology and imaging 10.64898/2026.06.25.26356599 medRxiv
Top 0.1%
18.5%
Show abstract

Accurate interpretation of planar ventilation-perfusion (V/Q) scintigraphy, used for diagnosing pulmonary embolism (PE) based on PIOPED/EANM guidelines, requires objective assessment of mismatched V/Q defects. Manual delineation of V/Q defects is time-consuming, subject to interobserver variability, and rarely performed in practice, limiting standardized reporting and quantification of disease burden. To address these challenges, we evaluated four modern AI models for automated segmentation of vascular perfusion defects in planar V/Q scans and compared their performance to human annotators. We retrospectively identified 2,118 patients who underwent planar V/Q scans at The Ottawa Hospital (June 2019-February 2023). Six standard projections (ANT, POST, LAO, RAO, LPO, RPO) were included. Four 2D neural networks (U-Net, nnU-Net, Swin UNETR, and a Bottleneck Transformer U-Net [BTU-Net]) were trained on 1,313 patients (7,878 projections) and validated on 329 (1,974 projections) using physician-annotated defects. A hold-out test set of 46 high probability patients was used to evaluate segmentation quality, and defect detection accuracy using free-response receiver operating characteristic (FROC) analysis, where BTU-Net was the only model performing on par with human readers, showing robust sensitivity across the entire range of segmentation probabilities. At 1.5 false positives per projection rate (FPPR), BTU-Net outperformed other models with a sensitivity of 0.529 {+/-} 0.026, On a separate hold-out set of low likelihood of disease patients (n=430), the lowest FPPR was 0.08 {+/-} 0.01 for BTU-Net (P<0.0001). BTU-Net enables rapid, consistent, and accurate interpretation of planar V/Q scans. Such tools may enhance diagnostic efficiency, standardize reporting, and support non-expert readers in evaluating PE.

6
Adaptive Post-Processing Recovers Most of the Gap to nnU-Net v2 in Head and Neck GTV Segmentation: A Paired Three-Arm HECKTOR 2025 Benchmark

Oyarzun Silva, R.; Hernandez Hernandez, P.

2026-08-31 radiology and imaging 10.64898/2026.08.28.26361649 medRxiv
Top 0.1%
18.2%
Show abstract

Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.

7
Automated Design of Patient-Specific 4D-Printed Phantoms for Quality Assurance of Adaptive Radiotherapy on a 1.5T MR-Linac

Hamkins, H. M.; Tam, K. H.; Sobremonte, A.; Jogi, S.; Koay, E.; Hassanzadeh, C.; Segars, P.; Tyagi, N.; Subashi, E.

2026-07-13 radiology and imaging 10.64898/2026.07.09.26357659 medRxiv
Top 0.1%
15.7%
Show abstract

Background: Independent end-to-end verification of adaptive radiotherapy on MR-Linac systems is limited by the lack of patient-specific phantoms able to reproduce imaging and dosimetric properties from CT and MRI scanners. We present a method for automated generation of 4D, patient-specific, multi-material 3D-printable phantoms for quality assurance of adaptive radiotherapy on a 1.5T MR-Linac. Methods: Patient images were automatically segmented using a pretrained deep learning model. The segmented structures were converted into high-resolution 3D meshes and assembled into printable phantoms. A dosimeter holder was inserted at user-defined anatomical locations, with orientation optimized to avoid traversal across heterogeneous tissue interfaces. Physiological motion was incorporated by generating phantoms from images at different timepoints and interpolating deformation fields to create continuous 4D models. Multi-material organs designed by mixing a set of six polymers at various proportions were used to reproduce tissue-specific imaging properties. The properties of material mixtures were evaluated in a clinical CT simulator and a 1.5T MR-Linac. Results: The proposed workflow enables automated generation of anatomically realistic phantoms with several types of embedded dosimeters. A discrete search method was designed for placement and immobilization of OSLD, film, and ion chamber dosimeters. Calibration curves for Hounsfield units were derived through variations in radiopaque material content, while MR signal intensity was modulated by gel and tissue matrix mixtures. Patient-derived abdominal phantoms were fabricated at multiple scales while replicating internal anatomical detail. Multi-dimensional phantom generation enabled continuous representation of motion states with consistent mesh topology across phases. Conclusions: We demonstrate an end-to-end workflow for automated generation of 4D patient-specific phantoms for MR-Linac quality assurance. The method combines realistic anatomy, embedded dosimetry, multimodal imaging properties, and physiological motion within a single fabrication framework. This approachmay enable an improved validation of adaptive radiotherapy workflows in MR-guided treatment devices.

8
Standardised evaluation and monitoring of site-specific AI performance with physical CT phantoms

Genske, U.; Laudani, A.; Yan, L.; Peng, Y.; Boening, G.; Ulas, S. T.; Wagner, M. P.; Diekhoff, T.; Hamm, B.; Jahnke, P.

2026-07-02 radiology and imaging 10.64898/2026.07.01.26357033 medRxiv
Top 0.1%
15.1%
Show abstract

Artificial intelligence (AI) applications in computed tomography (CT) imaging require objective and continuous testing, yet standardised methods for this purpose have not been established. Here, we present a framework using physical phantoms for standardised testing and monitoring of AI, demonstrated in liver lesion detection. We begin by designing phantoms tailored to the anatomical input domain expected by AI algorithms, and then systematically assess how AI performance is affected by variations in scanner technology and operation across two clinical CT systems. Next, we perform longitudinal monitoring, yielding consistent results over fifteen months on both systems. Finally, we validate clinical relevance by demonstrating that AI models trained on phantom data generalize effectively to patients and exhibit no evidence of phantom-specific adaptation. Our findings show that anatomically realistic phantoms enable standardised, site-specific testing and monitoring of AI, providing a proactive method for local and cross-institutional quality assurance.

9
Augmenting Deep Learning-Based PSMA PET/CT Metastasis Segmentation with a Population-Level Spatial Atlas

Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.

2026-08-31 radiology and imaging 10.64898/2026.08.26.26361439 medRxiv
Top 0.1%
14.8%
Show abstract

Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.

10
Correlation of OCT-Based Radiomic Signatures With Dose-Associated Radiation Response in Tumor Spheroids

Arndt, M. D.; Hansler, R.; Tirinato, L.; Tkachenko, A.; Seco, J.; Schepers, U.; Spadea, M. F.

2026-07-09 cancer biology 10.64898/2026.07.08.737210 medRxiv
Top 0.1%
13.3%
Show abstract

Background: Three-dimensional tumor spheroids are an established radiobiology model, but scalable, reproducible readouts of dose-dependent radiation response are lacking. We evaluated whether optical coherence tomography (OCT) radiomics can quantify dose-associated response in spheroids, and how it compares with conventional brightfield morphology. Methods: This in vitro, cross-sectional study used SAS oral squamous cell carcinoma spheroids seeded at two densities (5000 and 10000 cells), irradiated at 0 to 12 Gy, and imaged on days 1 to 11 post-irradiation. Each OCT acquisition yielded co-registered structural-intensity and speckle-variance volumes. Radiomic features (shape, first-order, texture) were extracted with Radiomics.jl, filtered for repeatability, correlation-pruned, and ensemble-ranked. Dose correlation was assessed by repeated 5-fold cross-validation across five regressors, comparing brightfield-only (BF), OCT-only, and combined OCT+BF feature sets with paired Wilcoxon tests. Results: OCT-only models consistently outperformed the BF baseline (median R2 0.77 to 0.85 versus 0.61 to 0.69; p<0.001 for all regressors). Adding brightfield to OCT gave no consistent benefit, reaching significance only for Random Forest (p=0.026, power 0.62). A compact shared feature subset combined brightfield area dynamics with OCT texture, shape, and speckle-variance descriptors, all showing low repeat-scan variability relative to cohort variability. Conclusions: OCT radiomics provides a sensitive, reproducible, label-free high-throughput readout of spheroid radiation dose response that outperforms the current brightfield-based approach, without requiring concurrent brightfield acquisition.

11
RadGuide AI: Development and Technical Evaluation of a General Nuclear Medicine Agent for Traceable Radiopharmaceutical Decision Support

Gu, X.; Zhu, H.; Zhong, F.; Teng, G.-J.

2026-07-10 radiology and imaging 10.64898/2026.07.09.26357614 medRxiv
Top 0.1%
13.2%
Show abstract

Background: Nuclear medicine and radiopharmaceutical development require coordinated radiochemistry, dosimetry, molecular imaging, radiation-safety and clinical decision processes. Current workflows remain fragmented, difficult to audit and poorly standardised for evaluating domain-specific AI support. Methods: We developed RadGuide AI, a nuclear medicine agent built around a traceable data-model-tool loop. Patent, literature and clinical-trial records were converted into 15,596 initial QA items; relevance screening, completeness checks, semantic deduplication and cross-validation retained 5,474 core QA items. MedGemma-27B-Instruct served as the foundation model and was adapted with LoRA. The system incorporated 55 MCP-wrapped tools covering radiopharmaceutical R&D, clinical decision support, imaging analysis and radiation-safety/dosimetry. Evaluation used a locked N=200 benchmark with predefined denominators, leakage control, expert scoring, statistical procedures, factuality audits and tool-execution metrics. Results: RadGuide-LLM achieved 88.5% answer accuracy (177/200; 95% CI, 83.3-92.2%) and a Macro-Average score of 21.5/25 (bootstrap 95% CI, 20.9-22.0), exceeding GPT-4o, DeepSeek-V3.2 and the base MedGemma model in this technical evaluation. Supplementary audits reported guideline compliance, terminology recall, knowledge coverage, tool-routing success and preclinical/phantom dosimetry agreement with explicit denominators and confidence intervals. Interpretation: RadGuide AI converts nuclear medicine queries into auditable retrieval, tool selection, calculation, verification and reporting workflows. The findings support technical feasibility, not definitive patient-level clinical validation; prospective multicentre studies and external benchmark release remain required before clinical deployment.

12
Photon-counting computed tomography for phantom-less quantitative measures of musculoskeletal tissues

Boyd, S. K.; Lackner, N. A.; Liphardt, A.-M.; May, M. S.; Schett, G.; Uder, M.; Engelke, K.

2026-07-23 radiology and imaging 10.64898/2026.07.22.26358681 medRxiv
Top 0.1%
13.1%
Show abstract

The advent of photon-counting computed tomography (PCCT) provides new opportunities to quantitatively measure musculoskeletal tissues such as bone, muscle and adipose because of the intrinsic use of spectral imaging. We aimed to evaluate the accuracy of measuring these tissues by PCCT under a range of scan protocols and compared our results to the current standard dual-energy CT (DECT). Phantoms containing inserts ranging from 50 to 200 mg/cm3 of calcium hydroxyapatite (HA) for estimating bone mineral density (BMD), and another phantom containing inserts for muscle and adipose tissues were scanned on PCCT and DECT at 120 and 140 kVp. We created virtual monoenergetic images (VMI) at energy levels from 40 keV to 190 keV for quantitative analyses. The averaged linear attenuation of phantom inserts was compared to theoretical values calculated from standardized attenuation profiles. Material decomposition using VMIs was compared to known HA concentration inserts to determine optimal image pairs for BMD measurement, notably without the need of phantom calibration. For most VMI energy levels the attenuation error was <1% for BMD at both 120 kVp and 140 kVp by PCCT compared to errors of <2% by DECT. The linear attenuation errors were <2.5% for muscle and <3.0% for adipose and results were similar for PCCT and DECT. Generally, errors were highest for low energy VMIs. Material decomposition using VMI pairs with a low energy at 50 or 60 keV and high energy between 150 and 190 keV produced calibration phantom-free estimates of BMD with <1% error. Results were similar for PCCT and DECT at 120 and 140 kVp. PCCT provides an accurate estimate of bone, muscle and adipose attenuation, and using material decomposition, estimations of BMD can be obtained without the need of phantom calibration.

13
LDCT-to-SDCT as a Bridge Problem: Single-Step Residual Endpoint Flow Matching for Real-Time Denoising

dela Sotta, T.; Saavedra, J. M.; Chang, V.; Xavier, A.; Henriquez, H.; Orellana, Y.; Curimil, J.

2026-08-31 radiology and imaging 10.64898/2026.08.27.26361520 medRxiv
Top 0.1%
13.0%
Show abstract

Diffusion models achieve high reconstruction quality in low-dose computed tomography (LDCT), but their iterative sampling trajectories impose substantial computational costs. Unlike unconditional generation, paired LDCT reconstruction starts from an image that already contains the anatomy and spatial structure of the standard-dose CT (SDCT) target; reconstruction primarily requires correcting dose-related noise and artifacts. We therefore introduce Residual Endpoint Flow Matching (REFM), an LDCT reconstruction method that learns to transport an LDCT image directly toward its paired SDCT endpoint rather than defining a noise-to-image trajectory. REFM predicts the residual velocity along linear interpolations between both images and supports single-step and multi-step reconstruction using the same trained network. We evaluate five model capacities using 1 to 50 Euler steps against deterministic U-Net and diffusion-based baselines. Across all REFM variants, one-step inference consistently provides the highest reconstruction quality. On the TCIA validation set, REFM Base achieves 50.98 dB PSNR and 0.9865 SSIM at 94.54 fps, compared with 50.92 dB, 0.9847, and 9.26 fps for DDPM-10. REFM Small retains 50.71 dB while increasing throughput to 198.56 fps. Without fine-tuning, REFM Base also matches the 25-step DDPM baseline on the external Mayo Clinic dataset, although DDPM remains stronger on synthetically degraded CRLM images. Thus, our results show that exploiting paired anatomical correspondence enables diffusion-level LDCT reconstruction with a single step reconstruction.

14
Assessment of MRI Geometric Distortion for Radiotherapy Planning

Odnovol, M.; Lykova, E.

2026-08-02 radiology and imaging 10.64898/2026.07.30.26359350 medRxiv
Top 0.1%
13.0%
Show abstract

Background: MRI is widely used in radiotherapy planning due to its high soft-tissue contrast, but geometric distortions can compromise target localization accuracy. Objective: This study aimed to develop an accessible method for assessing geometric distortion in MRI using two phantoms - a commercial anthropomorphic phantom and a custom-made phantom fabricated from ABS plastic. Approach: CT imaging was used as the reference standard. Distortion was assessed through linear measurements of periodic structures in a DICOM viewer, followed by statistical analysis. Significance: The study evaluates the clinical impact of distortion on radiotherapy planning and proposes a cost-effective solution for routine quality assurance in resource-limited settings.

15
Experimental hybrid spectral CT with Cramer-Rao lower bound-optimized weighting for quantitative iodine imaging

Sandvold, O. F.; Proksa, R.; Perkins, A. E.; Daerr, H.; Koehler, T.; Jacob, T.; Brown, K. M.; Roessl, E.; Noël, P. B.

2026-08-10 radiology and imaging 10.64898/2026.08.06.26359804 medRxiv
Top 0.1%
12.5%
Show abstract

Spectral computed tomography (CT) is a burgeoning quantitative imaging technique with applications in oncologic diagnostics, prognostic prediction, tissue perfusion studies, and treatment follow-up. While normalized iodine concentration values have been correlated with microenvironmental biophysical changes, obtaining accurate iodine concentrations, particularly at low concentrations remains difficult due to varying spectral CT instrumentation performance. Hybrid spectral CT systems, combining multiple spectral CT instrumentation techniques, address these quantitation insufficiencies by increasing spectral separation but have not been evaluated on a clinically analogous platform. We validate a hybrid spectral CT system, comprised of clinical-grade components, acquiring four distinct effective spectra and applying efficient noise-reducing weighting schemes to compare iodine noise and bias against conventional kVp-Switching (kVp-S). Two tube current levels (50, 350 mA) and three duty cycle ratios (33/67, 50/50, 75/25) were implemented to elucidate radiation dose exposure and kVp-S parameterization impact. A standard quality assurance (QA) and patient-derived, abdominal IodinePrint phantom were scanned on the system. The average absolute bias in iodine density images of the QA phantom was comparable across acquisition techniques, below 0.5 mg/mL, while quantitative noise improved by 22% using noise-optimized weighting schemes. In the IodinePrint phantom aorta and pancreas structures, the noise-optimized weighting scheme increased signal-to-noise ratio (SNR) by 1.3x compared to kVp-S alone. These results highlight the increased precision of hybrid, multi-channel spectral CT systems and motivate CT designs that enable robust CT biomarker development.

16
FreqFuseNet: Resolving Feature-Scale Mismatch in Dual-Frequency Fusion for Thin-Wall Head-and-Neck OAR Segmentation

Chen, W.-Y.; Wan, S.-Y.; Lin, G.-Y.

2026-07-13 radiology and imaging 10.64898/2026.07.09.26357642 medRxiv
Top 0.1%
12.0%
Show abstract

Accurate segmentation of thin-wall organs-at-risk (OARs)-the cochlea, vestibular semicircular canals, internal auditory canal, tympanic cavity, and middle ear-is clinically relevant for head-and-neck radiotherapy planning, yet these small, thin-wall structures remain among the most challenging targets for automated delineation. Dual-frequency feature fusion is a promising direction for boundary-sensitive representation, but under the investigated FP16 FFT-FcaNet setting, we observe an approximately 863-fold activation-scale mismatch between the FFT and FcaNet branches, causing a nominal 5 percent residual coefficient to behave as an approximately 43-fold dominant term. We propose FreqFuseNet, which resolves this mismatch by normalizing the FcaNet branch to the FFT activation scale before residual injection with a fixed low-amplitude coefficient (beta = 0.05), restoring beta as an interpretable 5 percent residual-amplitude coefficient relative to the FFT feature scale. Under a controlled binary per-OAR ROI protocol on the SegRap2023 head-and-neck CT benchmark across 10 clinically prioritized thin-wall OARs, FreqFuseNet achieves Dice of 0.849, HD95 of 0.824 mm, and SDice@1mm of 0.959 in the primary seed, with comparable performance in an independent second seed (Dice 0.843, HD95 0.823 mm). FreqFuseNet yields statistically significant case-level aggregate improvements over 3D U-Net and MedNeXt-S (Wilcoxon p < 0.01 and p < 0.05, respectively), using only 29.7 million parameters versus 414.6 million for the full wavelet baseline.

17
Feasibility of a 2-Minute Multi-Echo UTE Acquisition for Simultaneous CT-Like Bone-Weighted Imaging and Quantitative T2* Mapping of Short-T2 Tissue

Do, H. P.; Bekku, M.; Berkeley, D.; Golden, M.; Kitane, S.; Uike, M.; Shinoda, K.; Takayanagi, R.; Takai, H.; Kawai, T.; Seballos, K.; Conley, R.; Sorfleet, K.; Devries, D.; Tymkiw, B.; AlGhuraibawi, W.; Caruthers, S. D.; Kadbi, M.; Provencher, M.; Tashman, S.; Ho, C. P.

2026-08-19 radiology and imaging 10.64898/2026.08.18.26360232 medRxiv
Top 0.1%
11.7%
Show abstract

Purpose: To determine the feasibility of a 2-minute multi-echo UTE (mecho-UTE) for CT-like bone-weighted contrast and T2* quantification of tissues with short T2/T2*. Methods: Mecho-UTE data acquired from four patients and five healthy subjects were used to assess image quality of the CT-like contrast. All data were reconstructed using conventional gridding (GRID+CONV) and compared with those reconstructed using conjugate gradient SENSE combined with deep learning-based denoising (CG+DLR). Image resolution and sharpness of the CT-like images were assessed using the full width at half maximum (FWHM) and relative edge sharpness (RESH), respectively. Calimetrix UTE-T2* phantom was used to assess the accuracy of T2* quantification of the mecho-UTE sequence. Results: Two-minute mecho-UTE with CG+DLR has similar accuracy (0.37 {+/-} 0.27 vs. 0.67 {+/-} 0.54 ms, p=0.20) and better precision (0.28 {+/-} 0.16 vs. 1.23 {+/-} 0.29 ms, p<0.001) compared to the 5-minute mecho-UTE with GRID+CONV. The 2-minute mecho-UTE with CG+DLR has higher resolution and sharpness compared to the 5-minute scan with GRID+CONV. Conclusion: It is feasible to achieve simultaneous CT-like contrast and T2* quantification of short-T2 tissues in two minutes. When appropriately used, it may simplify logistics, reduce costs, and eliminate radiation exposure risks.

18
Dosimetric Characterization and Workflow Optimization of the FLASH-SARRP for Reliable Preclinical Radiobiological Studies

Knol, M.; Goncalves Jorge, P.; Kunz, L. V.; Korysko, P.; Petit, B.; Durham, A.; Marie-catherine, V.; Tsoutsou, P.; Koutsouvelis, N.; Lascaud, J.

2026-07-07 cancer biology 10.64898/2026.07.06.736680 medRxiv
Top 0.1%
10.8%
Show abstract

Objective: Preclinical small-animal irradiators such as the FLASH-SARRP can support the advancement of photon-FLASH toward the clinic. This study aimed at characterizing the FLASH-SARRP and established a robust quality assurance (QA) workflow to enable accurate and reproducible preclinical experiments. Approach: Custom 3D-printed spacers were designed to ensure reproducible X-ray tube alignment, sample positioning and mounting of the dosimetric tools. Beam characteristics were evaluated using a combined dosimetric approach. High spatially resolved dose distributions were obtained from Gafchromic films, whereas a plastic scintillating fiber was employed to monitor in real-time the temporal pulse structure and synchronization between the two X-ray tubes. Day-to-day variability of the delivery was evaluated over several sessions. Main results: The FLASH-SARRP achieved dose-rates of around 80 Gy/s when both tubes were used simultaneously and provided a homogeneous irradiation field suitable for small-animal studies. A desynchronization between the two tubes was observed with an average delay of 10 ms, resulting in temporal dose-rate heterogeneity. Additionally, a substantial inter-session variability (~11%) was found, whereas the intra-session variability was relatively low (~4%). Inter-session variability was reduced to 5%, approaching the intra-session variability, by adding Gafchromic films/scintillator-based quality assurance (QA) workflow into the irradiation routine. Significance: This work highlights the importance of temporal dosimetry for preclinical FLASH studies. Additionally, a practical QA framework is proposed integrating real-time monitoring with reference dosimetry. The proposed work enables adaptive dose delivery, thereby enhancing the reproducibility of the irradiations, which is crucial for reliable preclinical studies on the FLASH effect.

19
Modern Convolutional Design Improves Uterine MRI Segmentation, while nnU-Net Remains Most Robust Across Heterogeneous Datasets

Di Giovanni, D. A.; Takada, A.; McNabb, E.; Dana, J.; Yokota, H.; Tsuboyama, T.; Zakarian, R.; Vallieres, M.; Tsui, J. M. G.; Reinhold, C.

2026-07-24 radiology and imaging 10.64898/2026.07.22.26358583 medRxiv
Top 0.1%
10.4%
Show abstract

Purpose: To evaluate how segmentation architecture and dataset-adaptive configuration influence uterine MRI segmentation across heterogeneous benign and malignant tasks. Methods: U-Net, Swin-UNETR, and MedNeXt were compared with nnU-Net as a self-configuring reference across T2-weighted MRI datasets: public multiclass UMD anatomy/fibroid segmentation (n=300), institutional endometrial cancer tumor segmentation (n=206), and institutional uterine mass lesion segmentation (n=234). A relabeled external UMD-style cohort (n=12) assessed domain shift. Models used fixed partitions, fold ensembling, Dice, HD95, ASSD, volume error, and paired bootstrap comparisons with Holm correction. Results: MedNeXt was the strongest manually controlled architecture. nnU-Net achieved the highest performance on all internal datasets and external testing. Macro-Dice reached 0.761, 0.746, and 0.814 for nnU-Net on UMD, endometrial cancer, and uterine mass datasets, respectively, versus 0.722, 0.726, and 0.789 for MedNeXt. The nnU-Net-MedNeXt gap was largest for multiclass UMD segmentation and smaller in binary tasks. External testing degraded all models; nnU-Net remained highest (0.542), followed by MedNeXt (0.490), U-Net (0.396), and Swin-UNETR (0.287). Conclusions: Uterine MRI segmentation performance depended on task, architecture, and evaluation domain. MedNeXt supported modern convolutional design as a strong manual baseline, but nnU-Net remained the most robust overall, emphasizing the importance of dataset-adaptive configuration and external validation.

20
Accurate Hepatic Fat Fraction Quantification Across Body Sizes Using Photon-Counting CT

Li, X.; Kallman, C.; Zhang, D.; Guo, C.; Zhou, Y.

2026-07-30 radiology and imaging 10.64898/2026.07.28.26359152 medRxiv
Top 0.1%
10.0%
Show abstract

Objective: To identify common photon-counting CT (PCCT) virtual monochromatic imaging (VMI) settings for accurate hepatic fat fraction (FF) quantification across different body sizes, including large body habitus. Methods: Six non-iodinated fat lesions (FF 5%-40%) were embedded in anthropomorphic liver phantoms representing medium-sized (25x32.5 cm^2) and large (31x39 cm^2) abdomens. Phantoms were scanned on a PCCT system (NAEOTOM Alpha) at 120 and 140 kV. CT numbers were measured in VMIs at 40-190 keV in 1-keV increments. Linear regression between the ground-truth FF and measured Hounsfield units (HU) was used to estimate FF. Common optimal VMI settings yielding the minimum relative root-mean-square error (rRMSE) in both phantoms were identified. Results: A single VMI setting of 70 keV at 140 kV demonstrated the best overall performance across body sizes, with FF (%) = -0.689HU + 36.51 (R^2 > 0.996), achieving rRMSE [&le;]3.4% and absolute RMSE [&le;]0.7% in both phantoms. Robust performance (rRMSE [&le;] 5%) was consistently maintained across 69-71 keV using identical calibration parameters for both phantoms. These results represented a substantial improvement over previously reported dual-energy CT (DECT) performance, while enabling accurate quantification on PCCT at radiation doses approximately 40% lower than those used in prior DECT protocols. Conclusion: PCCT enables accurate and robust hepatic fat fraction quantification independent of body size. A single protocol at 140 kV with VMIs of 69-71 keV consistently achieved low quantification errors, demonstrating strong potential for opportunistic liver fat assessment using PCCT, especially in obese patients.